Observability for Large Language Models - Site Reliability and Chaos Engineering for AI at Scale
- Författare
- Ankush Sharma
- (Ankush Sharma., Part I: Foundations of Observability for LLMs.- Chapter 1: Introduction to LLMs and Observability.- Chapter 2: Site Reliability Engineering (SRE) Overview.- Chapter 3: Observability in AI vs. Traditional Systems.- Part II: Measuring Performance in LLMs.- Chapter 4: Defining Service Level Objectives (SLOs) for LLMs.- Chapter 5: Observability Metrics for LLMs.- Chapter 6: The Role of Logs in LLM Systems.- Chapter 7: Distributed Tracing for LLM Workflows.- Part III: Scaling Observability Across Distributed Systems.- Chapter 8: Observability in Multi-Model Enviroments.- Chapter 9: Capacity Planning and Scaling LLMs.- Chapter 10: Reducing Latency in LLM Systems.- Chapter 11: Fault Tolerant LLM Infrastructure.-Part IV: Chaos Engineering for LLM Reliability.- Chapter 12: Introduction to Chaos Engineering.- Chapter 13: Chaos Experiments for LLMs.- Chapter 14: Automating Chaos Engineering for AI.- Part V: Monitoring and Improving LLM Performance.- Chapter 15: Real-Time Monitoring Systems for LLMs.- Chapter 16: Postmortems for LLM Failure.- Chapter 17: Retraining and Model Drift Monitoring.- Part VI: AI Ethics and Accountability in Observability.- Chapter 18: Governance and Compliance in LLM Systems.- Chapter 19: Telemetry and Accountability.- Chapter 20: Future Trends in AI Observability.)
- Genre
- Facklitteratur
- Språk
- Engelska
| Förlag | År | Ort | Om boken | ISBN |
|---|---|---|---|---|
| Apress | 2026 | USA, Berkley | 236 sidor illustrationer | 979-8-8688-2826-3 |